Skip to content

Self-hosted model endpoints for any role; Hermes v2026.9.24 with reasoning echo - #438

Merged
renmengye merged 2 commits into
mainfrom
feat/self-hosted-endpoints
Sep 28, 2026
Merged

renmengye merged 2 commits into
mainfrom
feat/self-hosted-endpoints

Conversation

@renmengye

Copy link
Copy Markdown
Member

Operators want to run authors and judges on an open-weight model they serve themselves. vLLM exposes the OpenAI chat-completions, OpenAI Responses and Anthropic Messages APIs on one server behind a bearer key, but the kernel had no way to point a role at it.

Configuration

  • Named endpoint profiles: OUTERLOOP_ENDPOINT_<NAME>_URL (the OpenAI-style base, e.g. https://llm.example.internal/v1), _KEY_FILE (absolute, owner-only), _MODEL (the served model), _API (anthropic, responses or chat).
  • Roles select a profile with model[endpoint=name], or [endpoint=name] to use the profile's model: the author (OUTERLOOP_AUTHOR_ENDPOINT or the author model), panel lens specs, and standalone reviewers (REVIEW_ENDPOINT). The bracket syntax cannot collide with vendor ids that contain @, / or :.
  • Preflight rejects an unknown profile, a missing or permissive key file, or a backend/API mismatch before any spend.

Per backend

  • Claude Code: Anthropic API. The base URL drops a trailing /v1 (Claude Code appends /v1/messages), the auth token comes from the key file, and Vertex, Bedrock, Foundry and nonessential traffic are off, so a failing endpoint never falls back to a vendor.
  • Hermes: a named custom provider with key_env and model.reasoning_echo: true, so each turn keeps its earlier reasoning (without it, custom endpoints lose it every turn).
  • Codex: a custom model provider over the Responses API.
  • Keys travel only through the environment (APPTAINERENV_* in containers), are redacted from verdict data, and keep the existing judge/author key separation.
  • Run records now persist the author's route. Old records without it keep their native routing (legacy fixtures included).

Hermes pin: v2026.9.24. Audited against the new source: flags now use argparse (the quoting workaround is gone), reasoning_echo exists, and newer instruction-file names (AGENTS.override.md and aliases) are sanitized for judges. The trajectory format and custom-provider contract are unchanged.

Verified live against a self-hosted vLLM server, through the kernel's own build_harness/run_role, with a logging proxy:

  • Hermes: verdict filed; 13 of 13 follow-up requests carried earlier reasoning; all requests to the local chat-completions endpoint.
  • Claude Code: verdict filed; 14 of 14 carried earlier reasoning (thinking blocks); all to the local Messages endpoint. (The first live run found the doubled /v1, fixed here.)
  • Codex over Responses: requests reached the local endpoint and returned 200, but the session filed no verdict; the served model's tool calls do not appear to come back through that server's Responses API. Documented as a known limitation; Claude Code and Hermes are the supported backends for self-hosted endpoints until this is resolved.

Gate: 2352 passed, 2 skipped; ruff, format, mypy clean.

🤖 Generated with Claude Code

renmengye and others added 2 commits September 28, 2026 12:15
…oning echo

Operators can run authors and judges on a model they serve themselves.
Deployment settings define named endpoint profiles (OUTERLOOP_ENDPOINT_<NAME>_URL,
_KEY_FILE, _MODEL, _API); a role selects one with `model[endpoint=name]`, or
`[endpoint=name]` for the profile's model, in the author setting, panel lens
specs and REVIEW_ENDPOINT. Claude Code uses the Anthropic API (base URL without
/v1, auth token from the key file, vendor and cloud paths off), Codex a custom
Responses provider, Hermes a named custom provider with reasoning_echo so each
turn keeps its earlier reasoning. Keys travel only through the environment,
are redacted from verdicts, and keep the judge/author key separation. Old run
records keep their native routing.

Hermes moves to v2026.9.24: its flags are argparse now, reasoning_echo exists,
and newer instruction-file names are sanitized for judges.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing and the /v1 handling

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 6d5aa001 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

Advisory (non-blocking):

  • Endpoint reviewer settings never reach the shipped standalone reviewer workflows. [coverage+deployment] The endpoint resolver reads REVIEW_ENDPOINT and OUTERLOOP_ENDPOINT_<NAME>_URL, _KEY_FILE, _MODEL, and _API, but advisory-review-agent.yml exports none of them at this session environment and review-agent.yml neither defines inputs nor forwards them. As a result, hosted ubuntu-latest standalone reviews cannot select a profile and instead use the native path or skip. (.github/workflows/advisory-review-agent.yml:226; high confidence)

Merged one non-blocking configuration gap from the coverage and deployment lenses. Rejected findings: none; the credentials, general, lifecycle, and prose lenses reported no findings.

@renmengye

Copy link
Copy Markdown
Member Author

On the advisory: intended for now. Endpoint profiles target deployments on the same network as the model server; a GitHub-hosted runner cannot reach a private self-hosted endpoint, so the hosted review workflows keep the native path. Forwarding the profile settings through the reusable workflows is a follow-up for self-hosted runners.

@renmengye
renmengye merged commit 81a8155 into main Sep 28, 2026
1 check passed
@renmengye
renmengye deleted the feat/self-hosted-endpoints branch September 28, 2026 16:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant